Papers with analysis pipeline

2 papers
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)

Copied to clipboard

Challenge: Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday.
Approach: They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages.
Outcome: The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla.
Interpretable Semantic Gradients in SSD: A PCA Sweep Approach and a Case Study on AI Discourse (2026.findings-acl)

Copied to clipboard

Challenge: Supervised Semantic Differential (SSD) is a mixed quantitative–interpretive method that models how text meaning varies with continuous individual-difference variables . currently no systematic method exists for choosing the number of retained components, introducing avoidable researcher degrees of freedom in the analysis pipeline.
Approach: They propose a PCA sweep procedure that treats dimensionality selection as a joint criterion over representation capacity, gradient interpretability, and stability across nearby values of K.
Outcome: The proposed method is based on a corpus of short posts about artificial intelligence written by Prolific participants who also completed Admiration and Rivalry narcissism scales.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations